Papers with text-only methods
PatentVision: A multimodal method for drafting patent applications (2026.eacl-industry)
Copied to clipboard
| Challenge: | PatentVision integrates textual and visual inputs to generate patent specifications . existing systems fail to capture the nuanced interplay between textual, visual components . |
| Approach: | They propose a multimodal framework that integrates textual and visual inputs to generate patent specifications. |
| Outcome: | The proposed framework surpasses text-only methods in patent writing, the authors show . it integrates visual data to better represent intricate design features and functional connections . |
Cross-media Structured Common Space for Multimedia Event Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | We propose a new task to extract events and their arguments from multimedia documents . traditional methods target text, images or videos, but multimedia content is distributed via multimedia . |
| Approach: | They propose a method that encodes structured representations of semantic information from textual and visual data into a common embedding space. |
| Outcome: | The proposed method achieves 4.0% and 9.8% absolute gains on text event argument role labeling and visual event extraction. |
The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility Assessment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new dataset evaluates the credibility of statements made by Israeli public figures and politicians . a dataset of 1021 statements is used to assess the credibility and accuracy of statements . |
| Approach: | They propose a dataset to evaluate the credibility of statements by Israeli politicians . they use annotated statements manually annotating them for their credibility status . |
| Outcome: | The proposed model outperforms models based on statement and context, and achieves a 48.3 F1 score. |
Bridging the Sensory Gap: Visual Injection for Taxonomy Completion (2026.acl-long)
Copied to clipboard
| Challenge: | Existing text-only methods suffer from a "Sensory Gap" in integrating new concepts into existing hierarchies. |
| Approach: | They propose a framework leveraging Visual Injection for Taxonomy Completion that maps synthesized images into intrinsic pseudo-tokens and decouples magnitude from selection to prevent visual signals from being drowned out. |
| Outcome: | Experiments on three datasets show that VITC achieves state-of-the-art performance . it delivers an average absolute gain of over 19% in Hit@1. |
Point-of-Interest Type Prediction using Text and Images (2021.emnlp-main)
Copied to clipboard
| Challenge: | Prior efforts in POI type prediction focus on text without taking visual information into account. |
| Approach: | They propose to use multimodal information from text and images to infer the type of a place from where a social media post was shared. |
| Outcome: | The proposed method outperforms the state-of-the-art method for POI type prediction based on text-only methods and sheds light on cross-modal interactions and limitations. |
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models (2025.emnlp-main)
Copied to clipboard
Qihang Ma, Shengyu Li, Jie Tang, Dingkang Yang, null Chenshaodong, Yingyi Zhang, Chao Feng, Ran Jiao
| Challenge: | Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs. |
| Approach: | They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information. |
| Outcome: | The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests. |